Papers with decoder-based pre-trained language models
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing decoder-based pre-trained language models demonstrate excellent multilingual capabilities, but it is unclear how they handle multilingualism. |
| Approach: | They propose to examine the neuron-level internal behavior of decoder-based PLMs by finding neurons that fire “uniquely for each language” within decoded PLM models. |
| Outcome: | The proposed models fire “uniquely for each language” and show that language-specific neurons are unique, with a slight overlap (5%) between languages. |
What Matters in Memorizing and Recalling Facts? Multifaceted Benchmarks for Knowledge Probing in Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Language models often exhibit factual hallucination issue, exhibiting factual factual knowledge-grounded sentences. |
| Approach: | They introduce a knowledge probing benchmark to evaluate the knowledge recall ability of pre-trained language models from diverse perspectives. |
| Outcome: | The proposed benchmark evaluates the knowledge recall ability of encoder- and decoder-based pre-trained language models from diverse perspectives. |